Papers with multi-modal task of

2 papers
Learning the Effects of Physical Actions in a Multi-modal Environment (2023.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are trained on large corpora of disembodied texts.
Approach: They propose a multi-modal task of predicting the outcomes of actions solely from realistic sensory inputs (images and text). They extend an LLM to model latent representations of objects to better predict action outcomes in an environment.
Outcome: The proposed model can capture commonsense when augmented with visual information and generalize and learn commonsensical reasoning better.
Language-Driven Region Pointer Advancement for Controllable Image Captioning (2020.coling-main)

Copied to clipboard

Challenge: Controllable Image Captioning is a recent sub-task of Image Captions wherein constraints are placed on which regions in an image should be described in the generated natural language caption.
Approach: They propose a method for predicting the timing of region pointer advancement by treating the advancement step as a natural part of the language structure via a NEXT-token.
Outcome: The proposed method agrees with ground-truth timing in the Flickr30k Entities test data with a precision of 86.55% and a recall of 97.92%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations